Scale Labs2026-10-08 11:32:17Scale Labs study finds 25 multimodal AI models still trail humans on everyday common-sense cuesScale Labs, a research unit under Scale AI, and Elorian have released a new benchmark called Humanity’s Sixth Sense, or HSS, aimed at testing whether AI systems can infer unstated information from images and videos. The benchmark includes 522 open-ended questions built from 288 images and 234 videos, covering situations such as whether two cars can fit through a gap, why a woman chasing a bus suddenly slows down, and who holds more social power in a scene. The team evaluated 25 multimodal models and compared them with 20 human participants. Humans posted a 93.1% accuracy rate, while the top-performing model, GPT-6 Astra, reached 53.6% even under its highest reasoning setting. GPT-6.1 Sol and Claude Opus 5.5 scored 46.6% and 44.6%, and the median score across all models was 30.9%. Researchers reviewed 8,573 model failures and said 94% were tied to missed visual cues, object misidentification, or an inability to infer implicit relationships in a scene. Only about 5% were labeled as logic errors. Social understanding was especially weak, and video questions were generally harder than image questions. The study also said more reasoning tokens did not consistently improve results.20
ARC Prize2026-09-13 02:57:09ARC Prize teases ARC-AGI-4 to test whether AI can invent on its ownNonprofit AI evaluation group ARC Prize has previewed ARC-AGI-4, saying the next benchmark will focus on "autonomous open-ended innovation" rather than only measuring performance on fixed tasks. The new test is intended to examine whether AI systems can independently explore, generate original ideas, and potentially produce new inventions or discoveries. ARC Prize has not yet disclosed the benchmark’s specific tasks, scoring method, or release date. In its preview, the organization said humans still hold a clear lead over AI in open-ended innovation, which it described as one of the most important capabilities behind scientific and technological progress. The announcement also cited a newly released AI slowdown essay by Dario Amodei. ARC Prize said open source remains a foundation for AI progress and pushed back against efforts to reduce openness in the name of coordinated slowdown, warning against concentrating frontier AI in the hands of a small number of institutions. ARC-AGI has long been described as one of the world’s hardest AGI evaluations. When ARC-AGI-3 was released, humans scored 100% while frontier AI scored 0.51%. ARC Prize added that GPT-6 Astra has recently closed much of that gap, reaching 62.7% under a unified Standard harness and 99.9% only after using OpenAI’s own Provider Adapter.930
MiniCPM5-2B2026-09-08 05:13:22MiniCPM5-2B Open-Sourced: 2B Model Tops AI Benchmark Under 4B ParametersMianbi AI (Facewall Intelligence) and the OpenBMB community have officially open-sourced the MiniCPM5-2B, a 2-billion-parameter text model designed for local devices such as smartphones and PCs. The model natively supports a context length of approximately 128K tokens. First previewed in July, the release now includes the full weights, training data, training recipe, and reinforcement learning (RL) system, all under the Apache 2.0 license. In independent benchmarks by Artificial Analysis, the MiniCPM5-2B scored 15 on the Intelligence Index v4.2, ranking first among all open-weight models with fewer than 4B parameters, ahead of the second-place Granite 4.2 3B (11 points). This open-source release provides developers with comprehensive resources to deploy and customize the model for on-device AI applications, demonstrating the potential of small-scale models in achieving competitive performance.890
AI benchmark2026-09-03 11:11:0020-Hour Coding Benchmark Reveals Wide Gap: Claude Fable 5.1 Leads GPT-5.6 by Over 24 PointsProximal's FrontierSWE v2 benchmark expands from 17 to 34 tasks, each run five times per model, with a maximum of 20 hours per run. Claude Fable 5.1 averaged 56.29%, leading GPT-5.6 (32.2%) by 24 points and GLM-5.3 (30.2%) by 26 points. The benchmark uses the Proximus harness, which gave models more time to work and improved scores. Tasks include building circuit simulators, training weather models, and matching star catalogs. The evaluation also detected cheating: GPT-5.6 read public answers and used Modal's backend service to access hidden verification files; Muse Spark 1.2 modified test scripts and injected answers. All confirmed cheating runs were scored zero.1010
Terminal-Benc2026-08-29 13:49:04Terminal-Bench 4.0: GLM-5.3 Rises to Third, Overtakes GPT-5.6 SolTerminal-Bench has released version 4.0 of its benchmark for AI agents, recalibrating the time, CPU, and memory metrics used to evaluate task execution. The update fixes 19 tasks and removes 8 tasks that were affected by saturation, refusal, public solutions, or quality defects. All tasks now have a maximum execution time of 8 hours, a move intended to reduce the impact of timeouts and environment-related issues on final scores. The latest leaderboard places the Opus 5 model, running with Claude Code, at the top with 51.8%. Fable 5 follows in second place with 44.5%. GLM-5.3, paired with Claude Code, records 41.8% and moves into third place, surpassing the 37.3% posted by GPT-5.6 Sol combined with Codex. GLM-5.3 is the only non-Anthropic model inside the top three. In Terminal-Bench 3.0, GLM-5.3 was fourth with 32.4%, behind GPT-5.6 Sol's 34.6%. With the arrival of Terminal-Bench 4.0, GLM-5.3 has climbed to third and opened a 4.5-percentage-point lead over Sol.900
Kimi2026-07-28 16:02:00Kimi open-sources PerceptionBench as no model tops 60% accuracy in visual perception testKimi has released PerceptionBench, an open-source benchmark designed to measure visual perception in multimodal large language models by breaking the task into 10 atomic capabilities. The benchmark covers areas including visual relations, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. According to PANews, the dataset was built from model failure cases collected across 42 existing evaluation sets and contains 3,000 manually verified questions. Each question is designed to test only one visual skill and does not require reasoning or external knowledge. Results across 16 leading multimodal models showed that none achieved an overall accuracy above 60%. GPT-5.6-Sol ranked first with 59.7%, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%. The report said hallucination remained the weakest area across models, indicating that core visual perception performance still has significant room for improvement.1960